Back

Journal of Chemical Theory and Computation

American Chemical Society (ACS)

Preprints posted in the last 30 days, ranked by how well they match Journal of Chemical Theory and Computation's content profile, based on 140 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.

1
Toward Robust Characterization of Dynamic Binding Pockets: Lessons from the HBV Capsid Assembly Modulator Site

Perez-Segura, C.; Scott, L. W.; Zlotnick, A.; Hadden-Perilla, J. A.

2026-08-10 biophysics 10.64898/2026.08.06.743403 medRxiv
Top 0.1%
40.6%
Show abstract

Protein function often depends on ligand binding pockets that fluctuate among conformational states, altering their size, shape, topology, and accessibility, yet quantitative comparison of these dynamic cavities remains challenging because their boundaries are often inherently ambiguous. The measure volinterior algorithm uses fuzzy-boundary detection to characterize enclosed molecular spaces; here, the hepatitis B virus (HBV) capsid assembly modulator (CAM) binding site is used as a model system to develop and validate a practical workflow for applying the method to dynamic protein binding pockets. The resulting methodology provides practical guidance for parameter selection and evaluation, establishes a standardized protocol for quantitative characterization of the HBV CAM pocket, and demonstrates robust, reproducible performance across conformational ensembles derived from molecular dynamics (MD) simulations. More broadly, this work provides a reproducible strategy for adapting measure volinterior to other dynamic binding pockets, enabling consistent comparison of pocket geometry among independent structural studies. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/743403v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@d460b5org.highwire.dtl.DTLVardef@11915c2org.highwire.dtl.DTLVardef@1e37524org.highwire.dtl.DTLVardef@1fc1b7_HPS_FORMAT_FIGEXP M_FIG C_FIG

2
A two-bead-per-aminoacid coarse-grained MD model with hydrogen bonding (2BPA-HB) to probe DNAJB6b-mediated suppression of polyglutamine aggregation in Huntingtons disease

ADUPA, V.; Polet, J. D.; Dekker, M.; Onck, P. R.

2026-08-27 biophysics 10.64898/2026.08.24.746793 medRxiv
Top 0.1%
40.5%
Show abstract

Polyglutamine (polyQ) aggregation plays a central role in several neurodegenerative diseases, including Huntington's disease. DNAJB6b, a molecular chaperone involved in protein quality control, is known to efficiently suppress polyQ aggregation, but its anti-aggregation mechanism remains unclear. In this work we investigate the interaction between DNAJB6b and the polyQ region (Q48) of mutant Huntingtin Exon 1 (mHttEx1) using a custom-built coarse-grained molecular dynamics model. The model incorporates a two-bead-per-amino-acid representation with hydrogen bonding (termed 2BPA-HB), and is calibrated against all-atom molecular dynamics data in terms of geometry, hydrophobicity, and hydrogen bonding. The model reproduces the tertiary structure of DNAJB6b and its interactions with Q48, and reveals an inverse correlation between DNAJB6b concentration and Q48 aggregation propensity. Our simulations show that DNAJB6b co-condensates with polyQ molecules, thereby shielding the polyQ from forming the intermolecular hydrogen bonds necessary for amyloid formation. The 2BPA-HB CGMD model en- ables efficient exploration of DNAJB6b conformations, supporting future studies of chaperone-mediated aggregation suppression and therapeutic development.

3
Ab initio side-chain sampling with PUD+ enables high-fidelity protein dynamics across AI-driven and classical simulations

Wu, D.; Wang, T.

2026-08-11 biophysics 10.64898/2026.08.10.743906 medRxiv
Top 0.1%
39.3%
Show abstract

The fidelity of molecular dynamics (MD) simulations fundamentally depends on the quality and coverage of the ab initio data used to parameterize the underlying force field, yet the role of side-chain conformational space remains insufficiently explored. In this study, we systematically investigate how comprehensive ab initio sampling of dipeptide conformations--specifically targeting side-chain degrees of freedom--impacts force field accuracy and MD simulation predictive power. We present the Protein Unit Dataset Plus (PUD+), a 40-million-conformation quantum mechanical dataset featuring unprecedented coverage of both backbone and side-chain conformational space. Machine learning force fields trained on PUD+ and integrated into AI2BMD simulations demonstrate superior energy and force prediction accuracy, capturing high-fidelity protein folding dynamics and the conformational flexibility of long-side-chain systems. Furthermore, leveraging PUD+ to reparameterize the CMAP term of the classical ff19SB force field markedly improves the description of intrinsically disordered protein (IDP) dynamics and IDP-ligand binding. Collectively, these results demonstrate that ab initio sampling of dipeptide side-chain conformations enables high-fidelity modeling of protein dynamics across both AI-driven and classical simulation paradigms.

4
Towards transferable explicit-solvent coarse-grained models for biomolecular condensates

Toplek, F. B.; Borges-Araujo, L.; Lindorff-Larsen, K.; Everaers, R.; Souza, P. C. T.; Morozova, T. I.

2026-08-29 biophysics 10.64898/2026.08.27.747511 medRxiv
Top 0.1%
38.4%
Show abstract

Biomolecular condensates formed by intrinsically disordered proteins require molecular models that accurately describe proteins in both dilute solution and condensed phases. Explicit-solvent coarse-grained models offer an attractive balance between chemical resolution and computational efficiency. Yet, it remains unclear whether improving dilute-state properties is sufficient to obtain an accurate description of condensates. Here, we address this question by introducing minimal modifications to the Martini 3 force field that combine recent advances in bonded interactions with refined protein-water interactions and strengthened glycine self-interactions, while preserving the underlying chemical transferability of the model. The resulting model substantially improves the description of single-chain conformations across a diverse benchmark of disordered proteins. We then investigate phase separation of the well-characterized low-complexity domain of heterogeneous nuclear ribonucleoprotein A1 and its sequence variants. The model reproduces several key physicochemical properties of biomolecular condensates, including chain expansion in the dense phase, sequence-dependent intermolecular contacts, protein diffusion and its relation to single-chain dimensions, and hydration, while revealing quantitative limitations in condensate density, phase equilibria, and ion partitioning. Our results show that improving dilute-state behaviour translates into a better description of condensed-phase properties, including condensate density, but is not sufficient to quantitatively reproduce the equilibrium between the dilute and dense phases.

5
Development of force-field corrections for the RNA A-bulge motif

Kudo, T.; Ekimoto, T.; Yamane, T.; Ikeguchi, M.

2026-08-27 biophysics 10.64898/2026.08.26.747445 medRxiv
Top 0.1%
38.3%
Show abstract

Many functional RNA motifs adopt structures that deviate from the canonical A-form helix and are emerging targets for RNA-directed therapeutics. The microtubule-associated protein tau (MAPT) A-bulge motif (5'-GCAGU/5'-ACGU) is one such motif. Because its structure is stabilized by a delicate balance of local interactions, its accurate modeling remains a major challenge for molecular dynamics (MD) simulations. The experimentally determined nuclear magnetic resonance (NMR) structure of the MAPT A-bulge motif provides a stringent test of whether RNA force fields can accurately reproduce the experimentally observed conformation. Most current AMBER-family RNA force-field models have incorrectly favored a non-native base-triple state of the MAPT A-bulge motif over the experimentally observed stacked state. Structural comparison of the stacked and base-triple conformations revealed that overly favorable NH-N hydrogen bonds between the bulged adenosine and an adjacent Watson-Crick base pair were the primary source of this imbalance. We developed gHBfix-18Ab, an 18-component hydrogen-bond correction that distinguishes NH and NH2; donors. gHBfix-18Ab was combined with the previously developed OL3CP and NBfix0BPh corrections to generate the composite model gHBfix-18Ab*. This model restored the experimentally observed stacked state as the global minimum in the calculated free-energy profile and improved agreement with NMR-derived distance data for the A-bulge region. Importantly, gHBfix-18Ab* did not produce marked structural destabilization of the cUUCGg tetraloop, a widely used benchmark for RNA force-field validation, suggesting that the refinement preserves the stability of the unrelated RNA motif. These results demonstrate that targeted refinement of hydrogen-bond interactions provides a practical strategy for systematic improvement of RNA force fields toward more accurate modeling of noncanonical RNA motifs.

6
Coarse-grained models for simulations of double-stranded nucleic acids for mixed protein-nucleic acid condensates

Yasuda, I.; Tesei, G.; Yamamoto, E.; Yasuoka, K.; Lindorff-Larsen, K.

2026-08-20 biophysics 10.64898/2026.08.14.744942 medRxiv
Top 0.1%
38.3%
Show abstract

Biomolecular condensates function as membraneless compartments, and some protein condensates can selectively concentrate single-stranded nucleic acids while excluding double-stranded nucleic acids. Understanding how nucleic acid structure affects partitioning into condensates has important implications for nucleic acid activity and function within condensates. Here, we present a set of coarse-grained two-bead-per-nucleotide models for simulations of double-stranded RNA and DNA in the CALVADOS framework. Our models separately represent the backbone and base, and maintain the helical structures using an elastic network potential tuned to capture chain stiffness. For dsRNA, the base stickiness was tuned using experimental data on differential partitioning of single- and double-stranded RNA into Ddx4N1 condensates in order to account for reduced base accessibility upon duplex formation. This RNA structural selectivity varied with the balance of electrostatic and non-electrostatic interactions, as revealed by simulations of condensates of the CAPRIN1 disordered region at varying ionic concentrations and with an R-to-K sequence variant. Finally, we developed parameters for double-stranded DNA using a similar approach. We envision that the CALVADOS models for double-stranded RNA and DNA will be useful for studying co-condensates of proteins and structured nucleic acids.

7
Gaussian Accelerated Molecular Dynamics in GROMACS

Yang, Y.

2026-08-10 biochemistry 10.64898/2026.08.10.743837 medRxiv
Top 0.1%
35.4%
Show abstract

Gaussian accelerated molecular dynamics (GaMD) enhances conformational sampling by adding a smooth boost potential without requiring predefined collective variables, but an engine-integrated implementation has not been available in GROMACS. Here, we implement total-, dihedral-, and dual-boost GaMD in GROMACS 2025.4, including staged energy-statistics collection, GPU-based bias evaluation and force scaling, restart support, and outputs required for cumulant-based free-energy reweighting. The implementation was evaluated using four benchmark systems spanning conformational free energies, protein folding, and ligand recognition. For alanine dipeptide, a reweighted 100 ns GaMD trajectory recovered the major free-energy basins and rotational barriers in overall agreement with a 1000 ns conventional MD simulation. For chignolin and TC5b, all three independent trajectories for each system sampled native-like folded states from extended conformations within 300 ns and 1 s, respectively; the best TC5b structure had a minimum backbone RMSD of 0.03 nm from the experimental structure. In the benzene-T4 lysozyme system, two of five independent 500 ns trajectories captured both ligand binding and dissociation, yielding a bound pose with a minimum ligand RMSD of 0.06 nm from the crystal structure. Across all four systems, the boost-potential distributions were approximately Gaussian, and second-order cumulant reweighting resolved the expected conformational and binding free-energy basins. These results demonstrate that GROMACS-GaMD provides a practical, GPU-enabled, collective-variable-free enhanced-sampling framework for biomolecular free-energy calculations, protein folding, and ligand-binding studies.

8
Enhanced-Sampling Molecular Dynamics Recovers Rare Functional RNA Conformations Across Diverse Structural Contexts

Li, D.; Ken, M.

2026-08-24 biophysics 10.64898/2026.08.22.746453 medRxiv
Top 0.1%
26.1%
Show abstract

Accurate determination of RNA conformational ensembles is essential for understanding RNA function and advancing RNA-targeted drug discovery, yet lowly-populated alternative states remain difficult to resolve with atomistic detail. A central constraint is that experimental refinement can only select conformations already present in the starting library, making library generation the limiting step. Using the HIV-1 trans-activation response element (TAR) as a model system, we benchmarked conventional MD (cMD) against the enhanced sampling methods Gaussian-accelerated MD (GaMD), replica-exchange Gaussian-accelerated MD (Rex-GaMD), replica-exchange with solute tempering (REST2), and temperature replica-exchange MD (T-REMD), as well as the structure-prediction based methods FARFAR2 and AlphaFold 3. Each library was refined against experimental residual dipolar couplings (RDC) and validated independently using ensemble-averaged QM/MM chemical shifts. We showed that T-REMD produced the most accurate ensemble by both measures, and its advantage tracked with broader, more continuous coverage of the interhelical conformational landscape. Broad temperature-range T-REMD also sampled conformations resembling excited state 1 (ES1) and the U23-A27-U38 base-triple, without requiring these states to be specified during library generation. More accurate ensembles further improved coverage of experimentally observed ligand-bound TAR conformations and enhanced ensemble-based virtual screening, linking structural accuracy to functional utility. The same workflow applied to the preQ1 class I riboswitch and the UUCG tetraloop improved agreement with experimental data in both cases. Together, these results establish replica-exchange enhanced sampling, particularly T-REMD, as an effective strategy for constructing experimentally validated RNA ensembles and accessing conformations corresponding to rare functional substates.

9
RPDynaFlow: Generating RNA-Protein Conformational Ensembles by Atomic Conditional Flow Matching

Li, Y.; Lu, K.

2026-08-28 biophysics 10.64898/2026.08.28.747734 medRxiv
Top 0.1%
25.9%
Show abstract

Conformation ensembles of biomolecules provide the basis for understanding structural transformations and drug design. Deep-learning generative models have advanced protein and small molecule ensemble generation, while RNA-Protein complexes remain unaddressed due to the chemical heterogeneity, limited dataset size and the different flexibility scales of RNA and protein components. We present RPDynaFlow, a flow-matching model to generate conformation ensembles of RNA-protein complexes, trained on 600 ns trajectories of molecular dynamics(MD) simulation. The results show our model extends the sampling range of the phase space compared to MD simulation, which couldbe treated as a rapid and efficient complement to MD trajectoriesfor studying RNA-protein interactions.

10
Sequence-dependent conformational and mechanical landscapes of double-stranded nucleic acids

Sharma, R.; Patelli, A. S.; Singh, R.; Petkeviciute-Gerlach, D.; Gonzalez, O.; Maddocks, J. H.

2026-08-10 biophysics 10.64898/2026.08.09.740023 medRxiv
Top 0.1%
23.0%
Show abstract

The sequence-dependent mechanical landscapes of double-stranded nucleic acid (dsNA) remain largely unexplored beyond canonical dsDNA. We describe cgNA+, a coarse-grained predictive model of the mechanics of dsRNA, DNA:RNA hybrids, and epigenetically modified dsDNA, all parameterised from 1.26 milliseconds of atomistic simulations. cgNA+ predicts non-local sequence-dependent equilibrium shape and stiffness with errors an order of magnitude smaller than sequence-variability, while enabling exploration of numbers of sequences inaccessible to atomistic simulation. We show that dsNA equilibrium shape is strongly influenced by flanking sequence up to octamer context, with flexible dimer-steps more context-sensitive. CpG-modification alters equilibrium shape comparable to changes caused by single-nucleotide polymorphisms. Groove width analysis across dsNA decamers reveals strong sequence dependence, reflecting the differing characteristic helical geometry of dsDNA and dsRNA, whereas DRHs exhibit mixed behaviour depending on DNA-strand pyrimidine content. CTCF binding sites exhibit a distinct groove width signature. Persistence-length spectra from [~] 9 million sequences indicate that dsRNA is stiffer than dsDNA, whereas DRH exhibit intermediate stiffness modulated by DNA strand pyrimidine content. Persistence length increases upon CpG-modification, but decreases on hypermodification. Overall, the cgNA+ model enables a first, highly accurate, very large-scale, comparative study of sequence-dependent mechanics both within and across dsNA classes, demonstrating previously hidden regulatory layers. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=74 SRC="FIGDIR/small/740023v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@6de6acorg.highwire.dtl.DTLVardef@1435181org.highwire.dtl.DTLVardef@9c2b60org.highwire.dtl.DTLVardef@e3abf0_HPS_FORMAT_FIGEXP M_FIG C_FIG

11
Predictive all-atom simulations of disordered proteins and biomolecular condensates through osmometry-guided force-field optimization

Ivanovic, M. T.; von Roten, V.; Schuler, B.; Best, R. B.

2026-08-26 biophysics 10.64898/2026.08.25.747127 medRxiv
Top 0.1%
22.4%
Show abstract

All-atom simulations with explicit solvent provide the most detailed and accurate description of dynamics and mechanisms in intrinsically disordered proteins and their condensates. However, interactions involving charged residues and ions remain a persistent source of systematic error. Here we introduce an osmometry-guided optimization strategy that directly targets residue-residue, residue-ion and ion-ion interactions. Osmotic pressure provides key experimental information on molecular interactions and can be calculated directly and rapidly from simulations, enabling efficient iterative force-field optimization. The resulting parameters improve agreement of all-atom simulations with a range of experimental data: single-molecule FRET measurements for 16 monomeric intrinsically disordered regions; NMR relaxation data for a complex between an IDP and a folded protein domain; and mean FRET efficiencies and chain reconfiguration times of IDPs in biomolecular condensates of highly charged proteins. For such condensates, simulations with an osmometry-calibrated force field provide the missing link for predicting condensate dynamics across length and time scales. The presented optimization strategy is broadly extensible to other interaction classes, including those governing protein-DNA and protein-RNA assemblies.

12
Effect of Glycosylation on the Free Energy Landscape of the Catalytic Domain of Human Carbonic Anhydrase IX

Dey, R.; Mondal, D.; Chakraborty, D.; Taraphder, S.

2026-08-26 biophysics 10.64898/2026.08.25.747051 medRxiv
Top 0.1%
19.4%
Show abstract

N-linked glycosylation is known to modulate the catalytic function of human carbonic anhydrase (HCA) IX, yet its influence on the underlying free-energy landscape remains largely unexplored. In the present work, we combine extensive all-atom molecular dynamics simulations with kinetic transition network analysis to investigate the effect of glycosylation on the conformational organization of the catalytic domain of HCA IX in both monomeric and dimeric forms. The multidimensional conformational space is discretized into distinct free energy minima using the distribution of reciprocal interatomic distances (DRID), and the effective barriers separating them are estimated using the max flow-min cut formalism. The corresponding free energy landscapes are visualized in terms of disconnectivity graphs, which provide a faithful representation of underlying kinetics. Minimum free energy paths, mean first passage times, as well as frustration metrics are computed to further quantify the effect of glycosylation on landscape topography. Unglycosylated systems are found to exhibit predominantly funnel-like landscapes, with a limited number of metastable states in the vicinity of the native protein fold. In contrast, glycosylation enhances landscape complexity, resulting in a wide array of relaxation timescales. Strikingly, the two glycan chains affect the landscape topography in distinct ways, despite having closely matching sequences. Dimerization couples the glycan chain dynamics, with transitions between key metastable states involving coordinated motions of both the chains. Our work illustrates that interpretation in terms of disconnectivity graphs and transition networks could reveal important insights into the organization of glycoprotein energy landscapes.

13
Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Nadeem, H.; Kleiman, D. E.; Zhou, Y.; Leakey, A. D. B.; Shukla, D.

2026-08-18 biophysics 10.64898/2026.08.12.744498 medRxiv
Top 0.1%
19.0%
Show abstract

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

14
Transferable Collective Variable to accelerate Protein-Ligand (Un)Binding Transitions via Explainable Machine Learning and Intriguing Role of Ligand Solvation

Dhibar, S.; Jana, B.

2026-08-22 biophysics 10.64898/2026.08.21.746233 medRxiv
Top 0.1%
19.0%
Show abstract

The process of drug unbinding is of immense importance in the field of biophysics and therapeutics. The behavior of these systems is greatly influenced by their thermodynamic and kinetic properties. Therefore, it is crucial to accurately estimate the ligand binding free energies and rate of ligand dissociation, yet these processes are often governed by rare event transitions that lie beyond the reach of standard brute-force molecular dynamics simulations. While enhanced sampling simulations offer a solution, their efficacy is strictly contingent upon the selection of appropriate collective variables (CVs) which is non-trivial for complex systems like protein-ligand complexes. In this study, we present a method to derive optimized CV from transition state region (TS) via an interpretable machine learning (ML) model, Elastic Net. By employing some physically intuitive order parameters, the derived optimized CV from the TS-region greatly accelerate ligand binding-unbinding transitions and achieves rapid free energy surface (FES) convergence across diverse systems including buried and solvent exposed active sites such as Trpsin-benzamidine complex, host-guest systems and sodium epoxidase etc. Intriguingly significant contribution of the ligand hydration is found in the optimized CV which depicts crucial role of solvent in driving ligand binding-unbinding transitions. The estimated binding free energies for different protein-ligand complexes match quite well with experiments, while maintaining a low computational cost. The derived optimized CV is also used to calculate the ligand residence times across different systems and calculated residence times are within the experimental range for all systems, again with very little computational costs. Moreover, we show that the optimized CV constructed from TS region via an interpretable ML model is transferable across diverse systems, offering a robust and scalable framework for drug discovery and investigation of complex biomolecular recognition.

15
Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

Clifton, B. R.; Grieve, A. G.; Corey, R. A.

2026-08-09 biophysics 10.64898/2026.08.08.743655 medRxiv
Top 0.1%
18.7%
Show abstract

Proteins dynamically switch between a continuum of interconverting conformational states, and understanding these structural dynamics is important for understanding protein function and for developing therapeutics. Molecular dynamics (MD) simulations can provide insight into protein conformational ensembles, but sampling rare conformational states can require substantial computational resources. The recent development of AI-based approaches for generating protein conformational ensembles, such as the Biomolecular Emulator (BioEmu), offers a potential alternative, although it remains unclear whether these approaches can accurately capture the conformational landscapes, especially for special cases such as membrane proteins. Here, we assess the ability of BioEmu to model the conformational dynamics of a model membrane protein, the bacterial rhomboid intramembrane proteases GlpG. We find that BioEmu generates a range of conformations corresponding to both open and closed states of the rhomboid lateral gate, including states associated with different stages of the catalytic cycle. These conformations broadly correspond to states sampled during microsecond-timescale MD simulations, although BioEmu does not reproduce the full conformational landscape observed using MD. BioEmu also samples substantial conformational heterogeneity within the soluble domains of rhomboids, which are highly flexible and poorly represented in experimental structures. Overall, our findings demonstrate that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations. These results suggest that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

16
MEM-CALVADOS: A Residue-Level Model for Flexible Proteins at Membrane Interfaces

Saltutti, R.; Tesei, G.

2026-08-20 biophysics 10.64898/2026.08.12.743853 medRxiv
Top 0.1%
18.2%
Show abstract

Many membrane proteins contain intrinsically disordered regions (IDRs) that play key biological roles by providing structural plasticity, harboring sites for post-translational modifications, and mediating protein clustering and phase separation. Residue-level molecular models parameterized against experimental data have provided insights into how IDR sequence controls conformational properties and phase behavior in soluble proteins. Here, we extend this modeling framework to membrane-associated IDRs. We adapt a coarse-grained lipid model, iSoLFv2, for four phospholipids and combine it with CALVADOS, a residue-level model for IDRs and multi-domain proteins. Protein-lipid cross-interaction parameters are calibrated to reproduce predicted insertion and orientation of transmembrane proteins with diverse architectures. We validate the resulting model against Wimley-White free energies of transfer of hydrophobic peptides and against an experimentally refined conformational ensemble of a flexible membrane receptor. Finally, we show the applicability of the model to a membrane-associated assembly of signaling proteins. The model provides a computationally efficient framework for studying conformational ensembles and assembly of proteins at bilayer-water interfaces.

17
Interaction-Range Control of Synapsin Aggregation in a Coarse-Grained Model

Krott, L. B.; Puccinelli, T.; Oliveira, W. d.; Gomes, M. E. N.; Lomba, E.; Piazza, F.; Bordin, J. R.

2026-08-21 biophysics 10.64898/2026.08.16.745062 medRxiv
Top 0.1%
15.3%
Show abstract

Synapsin-1 is a multidomain neuronal protein containing extensive intrinsically disordered regions and is a key component of synaptic-vesicle condensates. Direct residue-level simulation of the collective organization of thousands of synapsin molecules remains computationally demanding. Here, we develop a coarse-grained description that connects residue-level CALVADOS 3 simulations to a one-particle-per-protein model. A potential of mean force between two synapsin molecules is obtained by umbrella sampling and represented by an isotropic effective interaction containing a short-range attractive region and a weak outer repulsive contribution. We compare two treatments of this interaction that differ only in the retention of the outer tail. Langevin dynamics simulations of effective proteins show aggregation upon cooling and compression in both models, but with markedly different collective organization. The shorter-ranged model progressively coarsens toward a single dense domain, whereas retaining the outer repulsive contribution favors the persistence of multiple mesoscale aggregates. The two models also display distinct relationships between aggregate size and particle mobility at low temperature. These results show that weak features of an effective protein-protein interaction can have pronounced consequences for collective synapsin organization at mesoscopic scales. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=95 SRC="FIGDIR/small/745062v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@854288org.highwire.dtl.DTLVardef@d2fc52org.highwire.dtl.DTLVardef@1b37d5aorg.highwire.dtl.DTLVardef@eac583_HPS_FORMAT_FIGEXP M_FIG C_FIG

18
Geometric characterization of the HSV - 1 glycoprotein B - amyloid β interaction in Alzheimer's disease using Forman-Ricci curvature

Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.

2026-08-29 bioinformatics 10.64898/2026.08.26.747308 medRxiv
Top 0.2%
13.2%
Show abstract

Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.

19
PMPNN-DDG: an accurate machine learning-based {triangleup}{triangleup}G prediction pipeline trained on a novel interpretable feature set extracted from ProteinMPNN

Jani, R.; Ahmed, S.

2026-08-27 bioinformatics 10.64898/2026.08.23.746499 medRxiv
Top 0.2%
13.1%
Show abstract

An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype-phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF -R = -0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.

20
Molecular dynamics descriptors for 1,079 post-translationally modified protein systems

Liu, K.; Qian, Q.; Peng, J.; Ma, D.; Yao, Y.; Zhao, J.; Chi, Y.

2026-08-24 biophysics 10.64898/2026.08.23.746532 medRxiv
Top 0.2%
12.6%
Show abstract

Databases of post-translational modifications (PTMs) catalogue modified sites and increasingly add static structural context, but trajectory-derived descriptors remain scattered across specialised tools and general molecular dynamics archives. Dyna-MO PTM brings together 1,079 AlphaFold 3-seeded systems covering lysine acetylation, lysine and arginine monomethylation, and serine, threonine and tyrosine phosphorylation. Each system is linked to three completed 10 ns replicas generated with CHARMM36m and TIP3P, for 32.37 s of aggregate sampling. A 118-column table joins simulation and quality-control provenance with global relaxation measures, site solvent exposure, rotamers, secondary structure and ionic-contact proxies. Versioned identifiers connect the records to starting structures, trajectories, manifests and analysis scripts. Researchers can use the resource to filter PTM contexts, reproduce descriptors, prioritise longer simulations and evaluate trajectory-analysis or generative methods. The trajectories describe finite-window relaxation rather than equilibrium free energies, kinetics or matched PTM effects.